Papers with judge prediction
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems (2025.naacl-long)
Copied to clipboard
| Challenge: | Traditional retrieval-augmented generation benchmarks use heuristics as the ground truth for evaluation, but require an expensive large language model (LLM) as a judge for a reliable evaluation. |
| Approach: | They propose to use large language models as a judge for retrieval-augmented generation benchmarks . they use heuristic metrics as input and a large language model as heuriistic input . |
| Outcome: | The proposed method couples heuristic features with large language models as judge for evaluation. |